Papers with qualitative metrics

4 papers
An Automatic Method to Estimate Correctness of RAG (2025.coling-industry)

Copied to clipboard

Challenge: Existing methods to assess the correctness of RAG models fail to capture the model’s internal state during answer generation.
Approach: They propose a method to predict the correctness of RAG models by modeling the model’s uncertainty on quantified perturbations of input.
Outcome: Extensive experiments across multiple large language models show that the proposed approach quantifies RAG robustness by aligning predictions with ground truth with a MSE 0.002 while offering flexibility for diverse qualitative metrics.
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for creating a vision question-answering with natural language explanations rely on human annotations that are time-consuming and costly.
Approach: They propose a method that generates high-quality natural language explanations using LVLMs by using visual prompts.
Outcome: The proposed method generates high-quality synthetic VQA-NLE datasets 20x faster than human annotations with minimal decrease in qualitative metrics.
Minimal Yet Big Impact: How AI Agent Back-channeling Enhances Conversational Engagement through Conversation Persistence and Context Richness (2024.findings-emnlp)

Copied to clipboard

Challenge: Increasing use of AI agents in conversational services highlights the importance of back-channeling (BC) as an active listening strategy to enhance conversational engagement.
Approach: They conducted an experiment with 55 participants to evaluate conversational engagement using both quantitative and qualitative metrics.
Outcome: The results show that the Todak_BC and TodAK_NoBC groups have significantly higher conversational engagement than the Todask_NoB.
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos (2025.emnlp-main)

Copied to clipboard

Challenge: This paper explores using Multimodal Large Language Models (MLLMs) to respond to student questions from online lectures . MLLM is a novel question answering task of real world significance .
Approach: They propose to use Multimodal Large Language Models to automatically respond to student questions from online lectures by using a dataset of 5252 question-answer pairs from 296 computer science videos.
Outcome: The proposed model can fine tune and fine tune questions from 296 computer science videos and show that students' preferences are important to the task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations